Back

Medical Decision Making

SAGE Publications

Preprints posted in the last 7 days, ranked by how well they match Medical Decision Making's content profile, based on 12 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.

2026-07-19 infectious diseases 10.64898/2026.07.17.26358308 medRxiv
Top 0.1%
2.4%
Show abstract

Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.

2
Rationale and guidance for implementing the continual reassessment method for dose-finding in controlled human infection model studies

Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.

2026-07-17 infectious diseases 10.64898/2026.07.16.26358128 medRxiv
Top 0.2%
1.7%
Show abstract

Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.

3
Multilevel Factors Associated with Nonresponse to Patient-Reported Outcome Measures in Routine Radiation Oncology Care

Liu, J. B.; Chen, Y.-J.; Edelen, M. O.; Pusic, A. L.; Martin, N. E.; Zeng, C.

2026-07-17 health systems and quality improvement 10.64898/2026.07.15.26358162 medRxiv
Top 0.3%
1.1%
Show abstract

Purpose: Nonresponse to routinely collected patient-reported outcome measures (PROMs) threatens the representativeness of aggregated data. We characterized patient-, provider-, and clinic-level factors associated with PROMIS Global-10 nonresponse in routine radiation oncology care. Methods: In this retrospective cohort study, all adults seen at five Mass General Brigham radiation oncology clinics over one year were included. The primary outcome was patient-level nonresponse, defined as never completing the portal-administered Global-10 versus completing it at least once. Using iterative mixed-effects logistic regression, we modeled patient-, provider-, and clinic-level factors. Results: Among 12,214 patients, 71 providers, and five clinics, patient- and appointment-level response rates were 35.4% and 10.9%, with patient-level response ranging nearly fivefold across clinics (12.8% to 66.2%). In Model 1, male sex, lower education, not working, and recent surgery had higher odds of nonresponse, and longer time since diagnosis lower odds. After provider- and clinic-level factors were added, patient sex, education, and employment became nonsignificant, whereas recent surgery (adjusted odds ratio [aOR] 1.97) and longer time since diagnosis (aOR 0.46 for >12 months) persisted. A provider's historical collection rate was protective but attenuated at the clinic level. There, a later program launch (aOR 0.29) and higher historical collection rate (aOR 0.79) correlated with lower nonresponse, whereas academic versus community setting did not. Conclusions: Nonresponse to routinely collected PROMs is a multilevel phenomenon driven substantially by clinic-level implementation factors, not patient characteristics alone. Because response rate is only a proxy for representativeness, PROMs programs and PRO-based performance measures should prioritize representative collection over volume.

4
Modeling effect of hypertension control on death, incidence of atrial fibrillation and economic impact to Medicare and hospitals.

Williams, J.; Mencer, N.; Mak, W. Y.; Dalle Luche, G.; Dundovic, S.

2026-07-17 health systems and quality improvement 10.64898/2026.07.15.26358198 medRxiv
Top 0.3%
1.1%
Show abstract

Background Hypertension is a major modifiable risk factor for atrial fibrillation (AF), yet blood pressure (BP) control remains suboptimal in older U.S. adults. Objectives This study evaluated how improve systolic BP (SBP) control could affect incident AF, downstream AF ablation demand, Medicare savings, and hospital revenue. Methods A population-based modelling framework was developed to estimate mortality and incident AF hazards across SBP strata: <120, 120-139, 140-159, and ?160 mm/Hg. AF incidence in the SBP <120 mmHg group was set at 2.2 per 1,000 person-year, with hazard ratios of 1.17, 1.42 and 1.64 applied to higher SBP strata. We assumed 25% of incident AF patients would undergo ablation, with a 7.2% complication rate. AF prevalence was projected to increase by 4.6% annually over 10 years. Medicare savings and hospital revenue foregone were estimated under varying procedure cost and contribution-margin assumptions. Results Higher SBP was associated with greater hazards of death and incident AF. Improved SBP control reduced projected AF incidence and ablation demand. Over 10 years, cumulative Medicare savings were projected at $8.7B-$10.9B across the full modelled population. However, reduced ablation volume translated into hospital revenue foregone, ranging from $75M to $377M in the first year, and approximately $1.03B-$5.2B cumulatively over 10 years. Conclusions Improved SBP control may reduce AF incidence, prevent avoidable invasive ablation procedures, relieve pressure on surgical waitlists, and generate substantial Medicare savings. However, these benefits may reduce hospital procedural revenue, highlighting a misalignment between prevention-oriented care and fee-for-service reimbursement incentives.

5
Selective prediction as a triage gate for primary-care depression screening: quantifying and mitigating selection bias in CHARLS-2011

Wang, Z.; liu, y.

2026-07-20 health informatics 10.64898/2026.07.17.26357845 medRxiv
Top 0.3%
1.0%
Show abstract

Background Primary care in China lacks structured mental-health assessment, and the machine-learning models that could support such screening are typically developed on heavily selected samples. Cumulative inclusion and exclusion criteria, though usually treated as neutral data-cleaning steps, can create heterogeneity in predictive reliability among retained participants. Using the China Health and Retirement Longitudinal Study (CHARLS) 2011 baseline, we quantified how selection funnels distort epidemiological associations and inflate machine-learning metrics, and tested selective prediction as mitigation. Methods Using the CHARLS 2011 baseline with temporal external validation in CHARLS-2018, we built a four-level selection funnel (L0-L3), evaluated five classifiers with nested cross-validation and SMOTE, and compared model-embedded uncertainty with a decoupled predictor-selector framework; XGBoost cross-validation residuals drove risk stratification and classification and regression tree (CART) rules. Results Sample sizes fell from L0 n=17,705 to L3 n=4,256 (24.0%). The cancer-depression odds ratio attenuated from 1.78 (95% CI 1.32-2.41) to 1.39 (0.74-2.63), losing significance. AUC rose with selection but not after multiple-comparison correction, whereas calibration error increased for four of five models. Model-embedded uncertainty succeeded only for XGBoost; with the decoupled XGBoost residual selector, all five models achieved selective prediction at approximately 20% coverage (test AUC 0.90, 95% CI 0.85-0.95), abstaining on approximately 80% of cases for individual safety. Risk stratification was stable (residual Spearman correlations >0.95; multi-seed Jaccard 0.88), and CART rules used self-rated health, education, pain, and marital status. Conclusions The findings support a deployable primary-care triage pathway: a four-variable rule identifies patients suitable for algorithm-assisted scoring (approximately 20% coverage) and routes the remainder to human evaluation. Methodologically, cumulative selection bias produces a dual distortion: epidemiological associations are compressed and machine-learning metrics inflated. Selective prediction is limited mainly by uncertainty-indicator design. Performance metrics should be reported with selection level, coverage, and calibration trajectory. Decoupled selective prediction with CART rule extraction provides an actionable framework for quality-controlled, tiered-care deployment. Keywords: selective prediction, selection bias, CHARLS, depression, predictor-selector decoupling, uncertainty quantification, classification and regression tree, triage, clinical decision support, health management.

6
What Week 8 Knows: Forecasting Six-Month GLP-1 Outcomes

Erly, B.; Raja, S.

2026-07-16 health informatics 10.64898/2026.07.14.26357506 medRxiv
Top 0.3%
0.9%
Show abstract

Background. Patients on GLP-1 medications lose very different amounts of weight, and most published prediction models include only patients who complete six months. That design omits everyone who disengages earlier, which is the majority of the cohort. We built a tool that includes patients who disengage and delivers useful predictions at the week-8 visit, where the clinical decision is actually made. Methods. Beginning with 237,800 adults enrolled in a US telehealth GLP-1 program, we required a documented week-8 weight, a refill-confirmed dose, and reported ethnicity, yielding an analytic cohort of 22,538. We answered three questions: the patient's likely six-month weight loss and our confidence in it; the probability of dropout before six months; and when weight loss plateaus. For the first, we fit a cubic in week-8 percent loss plus 16 covariates, with quantile-regression bands at the 10th and 90th percentiles for the prediction interval, checking fractional-logit and isotonic recalibration as alternatives. For the second, we fit a logistic regression and compared it to gradient boosting. For the third, we fit a per-patient exponential trajectory among patients with at least four weight observations. We trained on enrollments before 2024-07-01 and tested on later ones, compared completer outcomes to published RCTs, and tested the week-8 anchor against measurements at weeks 2, 4, 6, 8, 10, 12, 16, and 20. Results. Mean six-month weight loss in completers was 11.7% on semaglutide and 14.1% on tirzepatide, in line with STEP-1 and SURMOUNT-1. Six-month disengagement was 66%. The prediction model reached test R2 = 0.65 with a mean absolute error of 2.76 percentage points. Calibration was strong: calibration-in-the-large was -0.52 pp and the calibration slope was 0.96. The 80% quantile-regression interval covered 76% of test patients; the 95% interval covered 93%. The disengagement model reached test AUC 0.79, against 0.74 for gradient boosting. Median plateau time among engaged patients was 387 days, longer in lower-BMI tertiles. The week-8 anchor gave R2 = 0.65, compared to 0.48 to 0.61 at earlier weeks and 0.67 to 0.91 at later weeks. We chose week 8 because 80% of slow responders reach their post-titration decision point at or before that visit. Two of twenty subgroup cells had reduced predictive accuracy; two more were too sparse to validate. Conclusions. Observed week-8 weight loss is the strongest predictor of six-month outcome. The model's accuracy (R2 = 0.65, MAE 2.76 pp) is appropriate for calibrating expectations and identifying patients for the post-titration decision, but not precise enough to drive that decision on its own. Disengagement is predictable at week 8 with AUC 0.79. Engaged patients plateau at a median of 387 days. Week 8 is the earliest visit at which titration is mostly complete, accuracy is in a useful range, and the post-titration decision remains actionable; later anchors predict better but inform a decision that has already been made for most patients. The model is temporally (internally) validated but not yet externally validated, and because it was developed on a single platform it should be regarded as a recalibration target rather than a drop-in deployment elsewhere. The tool is published as a public web calculator to support shared decision-making, though it is not precise enough on its own to drive an irreversible clinical decision. It is prognostic, not therapeutic; treatment-effect estimation is addressed in companion work.

7
Efficient stochastic epidemic simulation via the Sellke construction

van Boven, M.; Bootsma, M. C.

2026-07-17 epidemiology 10.64898/2026.07.16.26358219 medRxiv
Top 0.4%
0.9%
Show abstract

Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.

8
Optimizing Latent Tuberculosis Treatment Strategies Among Immigrants From High-Burden Settings

Tabackman, A.; Karoly, M.; Jacobson, K.; Horsburgh, C. R.; Linas, B.; Campbell, J.; Acuna-Villaorduna, C.; Sinha, P.

2026-07-16 health economics 10.64898/2026.07.13.26357974 medRxiv
Top 0.5%
0.6%
Show abstract

Importance Tuberculosis preventive therapy is central to reducing tuberculosis, and foreign-born individuals account for most US tuberculosis cases. Current US Preventive Services Task Force guidance recommends testing and treating all foreign-born individuals regardless of age or time since immigration, yet the risks of disease progression and of treatment-related harm are not uniform across these groups. Objective To evaluate the cost-effectiveness and health outcomes of tuberculosis infection treatment strategies among immigrants from high-burden settings, stratified by age and time since immigration. Design Decision analytical model using individual-level microsimulation (Markov model) over a 30-year horizon, with deterministic and probabilistic (second-order Monte Carlo) sensitivity analyses. Costs and outcomes were discounted at 3%. Setting US federally funded tuberculosis clinic care (healthcare-sector perspective), using observed data from the Boston Medical Center/Boston Public Health Commission tuberculosis clinic and published literature. Participants A simulated cohort of 10000 IGRA-positive, foreign-born adults from high tuberculosis incidence settings (excluding immunosuppressed individuals), modeled as recent or remote (immigrated 25 years earlier) immigrants at ages 35 and 65 years. Interventions Rifampin daily for 4 months, isoniazid daily for 9 months, or no preventive therapy. Main Outcomes and Measures Costs, disability-adjusted life-years (DALYs), incident tuberculosis cases and deaths, treatment completion, and incremental cost-effectiveness ratios (ICERs), with the proportion of simulations in which each strategy was optimal at a willingness-to-pay threshold of $50000 per DALY averted. Results Among recent immigrants, rifampin was the dominant strategy at ages 35 and 65 years (optimal in 88.5% and 93.9% of simulations), yielding the fewest tuberculosis cases (119.44 and 82.31 per 10 000) and the highest treatment completion (71.4% and 67.7%). Among remote immigrants, rifampin remained the dominant strategy (optimal in 53.41% of simulations), followed by no treatment. In older remote immigrants, no treatment was optimal in 94.7% of simulations. ICERs for treatment vs no treatment were unfavorable ($193 600 and $412 857 per DALY averted for rifampin and isoniazid, respectively, at age 65). Conclusions and Relevance In this decision analytical model, rifampin was cost-effective for recent immigrants, whereas no treatment was optimal for older remote immigrants. Age and time since immigration may help risk-stratify tuberculosis infection treatment and reduce unnecessary treatment in lower-risk populations.

9
Projected burden of hypertension-associated cardiovascular disease in people living with HIV versus HIV-negative adults in Eswatini

Milali, M. P.; Citron, D. T.; Bhamidipati, K.; Yamamoto, N.; Osei-Ntansah, A.; Platais, I.; Ferrara, G.; Ngcamphalala, C.; Dlamini, S. G.; Ginindza, N.; Bershteyn, A.

2026-07-20 public and global health 10.64898/2026.07.17.26358301 medRxiv
Top 0.5%
0.6%
Show abstract

Background Having achieved the UNAIDS 95-95-95 targets, Eswatini faces a growing burden of non-communicable diseases which are major contributors to morbidity and mortality. Hypertension-associated cardiovascular disease (CVD) is rising among people living with HIV (PLHIV) as survival improves and metabolic risks - including those linked to dolutegravir (DTG) - increase. We projected CVD burden among PLHIV and HIV-negative adults (PLWHIV) through 2045 to inform integrated HIV-CVD planning. Methods EMOD-HIV, an agent-based model calibrated to Eswatini's epidemic, generated HIV prevalence trajectories. These were combined with age-standardized Global Burden of Disease CVD estimates and published relative risks (RRs) to produce HIV-stratified CVD projections. CVD burden trajectories were then projected through 2045 using a logistic generalized additive model with Monte Carlo uncertainty quantification. Five scenarios were evaluated to assess how different assumptions about RR of CVD among PLHIV versus PLWHIV affect projected burden: (1) CVD prevalence under a constant RR; (2) CVD mortality under a constant RR; (3) HTN-attributable CVD mortality under a constant RR; (4) HTN-attributable CVD mortality under a post-DTG RR increase following Eswatinis 2021 dolutegravir rollout; and (5) HTN-attributable CVD mortality under a gradual RR increase from 2010-2045 reflecting cumulative metabolic and demographic shifts. Results PLHIV consistently exhibited higher CVD burden than HIV-negative adults. Scenario 1: CVD prevalence was 11.0% (95% UI: 9.0-13.5%) among PLHIV versus 6.8% (6.0-7.8%) in 2025, stable through 2045. Scenario 2: CVD mortality rate was 0.60% (0.47-0.76%) versus 0.37% (0.31-0.45%) in 2025, declining modestly through 2045 with consistent excess. Scenario 3: HTN-attributable mortality was 73% (70-76%) versus 64% (61-66%) in women and 61% (58-64%) versus 58% (55-60%) in men, stable through 2045. Scenario 4: Following DTG rollout, mortality rose from 73% to 85% in women and 61% to 70% in men by 2022, remaining stable thereafter. Scenario 5: By 2045, mortality reached 85% (82-88%) in women and 67% (64-70%) in men with HIV, versus 60% (57 - 63%) and 55% (53 - 57%) in HIV-negative adults. Conclusions While excess CVD burden among PLHIV is projected to persist even under stable risk conditions, ART-related metabolic trajectories - particularly those linked to DTG - may drive substantial widening of this gap through 2045. HTN-attributable CVD mortality is particularly elevated among women with HIV. Strengthening integrated HIV-NCD services, including blood pressure screening, risk-based therapy, and sex-specific DTG counseling, will be essential to sustain long-term health gains.

10
Inequalities in Colorectal Cancer Screening: Combining MAIHDA with Difference-in-Differences to Assess Programme Effects Across Population Subgroups

Jolidon, V.; Delaruelle, K.; Kawachi, I.; Cullati, S.; Bell, A.; Holman, D.

2026-07-18 health policy 10.64898/2026.07.16.26358242 medRxiv
Top 0.5%
0.6%
Show abstract

Background: Research consistently shows that colorectal cancer (CRC) screening uptake is socially patterned; however, sociodemographic determinants are usually analysed separately, overlooking how multiple social conditions jointly shape inequalities. This also applies to policy research, where heterogeneity in screening programme effects remains underexplored. Methods: Using data from the European Health Interview Survey (2014 and 2019; n=201,214; 24 countries), we applied Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) to analyse CRC screening uptake across 72 subgroups defined by sex, education, living arrangement and employment. To assess heterogeneity in screening programme effects, we combined MAIHDA with difference-in-differences (MAIHDA-DiD). Results: MAIHDA revealed inequalities in uptake: lower- and middle-educated men, whether employed or unemployed, had the lowest uptake, whereas men and women not living alone, retired or living with disability, had the highest uptake. Lower-educated homemaker women were the only female group with below-average uptake. MAIHDA-DiD showed that programmes increased overall uptake but did not produce larger gains among groups with lower pre-intervention uptake, and therefore did not reduce inequalities. Instead, programmes generated above-average increases among groups with higher pre-intervention uptake, particularly lower- and middle-educated men and women not living alone and retired. Living arrangement explained more variation in programme effects than other factors, with individuals living alone benefiting less from the programmes. Conclusion: CRC programmes did not reduce (and may have widened) inequalities, underscoring the need for equity-focused strategies in population-based screening. By extending MAIHDA with difference-in-differences, this study introduces a novel approach for evaluating heterogeneous policy effects in public health.

11
A New Method to Predict the Effect of an Intervention in the Host Population to Reduce the Magnitude of an Outbreak of a Vector-Borne Infection

Coutinho, F. A. B.; Amaku, M.; Kallas, E. G.; Massad, E.

2026-07-19 epidemiology 10.64898/2026.07.16.26358272 medRxiv
Top 0.6%
0.5%
Show abstract

In this paper, we propose a new model to estimate the impact of an intervention on human hosts of a vector-borne infection, such as dengue, which occurs in yearly outbreaks of different magnitudes. The model applies to these outbreaks and, in fact, is independent of their intensity, that is, it does not require the steady-state assumption. The model takes as input the officially reported age-dependent number of cases of a vector-borne infection. It is deterministic and does not account for stochasticity. Our objective is to estimate the impact of the intervention (the efficacy), and we rely on the observed fact that the age distribution of the proportion of cases of the infections transmitted by the same vector is independent of both the intensity of transmission and the geographic area studied, at least for Brazilian regions. This finding is highlighted in the main text and forms the basis of our calculations. A hypothetical intervention is simulated using a dengue vaccine, which allows the determination of the optimal strategy for a vaccination campaign.

12
Impact of subgroup classification accuracy on detecting heterogeneous treatment effects in Staphylococcus aureus bacteraemia: A simulation study

Hamilton, F. W.; Ong, S. Y.; Swets, M.; Russell, C. D.; Underwood, J.

2026-07-20 infectious diseases 10.64898/2026.07.17.26357924 medRxiv
Top 0.7%
0.5%
Show abstract

Background Staphylococcus aureus bacteraemia (SAB) is clinically heterogeneous. Potential heterogeneous treatment effects (HTE) have recently been identified through analysis of patient subgroups, identified using routine clinical variables.However, the impact of misclassifying patients into these groups is unclear, and practical strategies to improve HTE detection remain uncertain. Methods We performed a simulation study using published data from selected randomised trials and observational studies in SAB. We assessed the impact of varying classification accuracy (70%-100%) on i) power, ii) type I error, and iii) bias in post-hoc analyses of HTE. We then evaluated two strategies to improve performance: enrichment designs, in which only patients predicted to belong to a target subgroup are randomised, and the use of ordinal rather than binary outcomes. Results Even with perfect classification, post-hoc detection of heterogeneous treatment effects remained highly conditional on subgroup prevalence, baseline mortality, and effect size. One subgroup was detectable at moderate sample sizes; however, power was inadequate for all other subgroups even with sample sizes of 20,000. Decreasing classification accuracy reduced power, increased type I error, and introduced bias. Enrichment marginally improved power. Ordinal outcomes substantially improved performance when they matched the treatment-effect structure, but were worse when they did not. Conclusions Detecting HTE in SAB is challenging, but not uniformly infeasible. Feasibility depends on the interaction between subgroup frequency, baseline risk, classifier performance, and outcome choice. To advance stratified medicine in SAB, research should prioritize robust classifiers, outcome measures matched to the expected mechanism and pattern of treatment effect, and trial designs that acknowledge uncertainty in subgroup prevalence and treatment-effect structure.

13
ReCo: a self-configuring and self-extending agentic framework for biomedical research

Tzanis, E.; Klontzas, M. E.

2026-07-16 health informatics 10.64898/2026.07.14.26358025 medRxiv
Top 0.8%
0.4%
Show abstract

This study presents ReCo (Research Cosmos), a self-configuring and self-extending agentic research framework for the biomedical domain. ReCo is orchestrated by a large language model that interacts with native computing tools, bundled Model Context Protocol (MCP) servers, structured skills, persistent project memory, and a desktop interface. Its bundled MCP servers provide biomedical analysis capabilities while serving as implementation paradigms for integrating new computational and AI frameworks. Structured skills encode procedures for environment configuration and framework ingestion, enabling ReCo to inspect repositories, manuscripts, or local codebases; identify dependencies and execution patterns; create isolated runtime environments; design and implement MCP interfaces. Self-extension was evaluated using five heterogeneous systems: the Merlin computed tomography foundation model, MAISI-v2 medical image synthesis framework, asari liquid chromatography-mass spectrometry workflow, DosimeTron agentic radiation-dosimetry platform, and Orthanc DICOM server. ReCo successfully operationalized all five systems and completed predefined functional evaluations. Re-hosted DosimeTron outputs demonstrated near-perfect agreement with the reference pipeline across 651 organ observations (Pearson correlation and Lin concordance correlation coefficient, 0.99999; mean absolute percentage difference, 0.37%). Notably, ReCo configured Orthanc as a PACS-like coordination layer, integrated it with DosimeTron, Merlin, and TotalSegmentator, and orchestrated data retrieval, analysis, and return of valid DICOM RTSTRUCT, RTDOSE, and Structured Report. ReCo provides a unified environment for configuring, documenting, and operationalizing heterogeneous biomedical frameworks, reducing technical barriers to the adoption and integration of emerging computational and AI methods. The official open-source ReCo GitHub repository is available at: https://github.com/eltzanis/ReCo

14
Design tensions in a two-sided marketplace for reusable digital therapeutics software components: a qualitative interview study

Kowatsch, T.; Melamed, S.; Nissen, M.; Merz, Y.

2026-07-20 health informatics 10.64898/2026.07.17.26358332 medRxiv
Top 0.8%
0.4%
Show abstract

Objectives To identify stakeholder-perceived design tensions in a two-sided marketplace for reusable digital therapeutics (DTx) software components and to use these tensions to propose alternative marketplace concepts. Methods We conducted 24 semi-structured interviews with digital health researchers and professionals. Data were analysed using hybrid deductive-inductive codebook thematic analysis. The Magic Triangle provided the initial deductive structure. One researcher coded all transcripts; a second independently applied the developing codebook to five transcripts to refine definitions and consistency. Seventeen parent themes were synthesized into 12 design tensions, which informed three author-generated marketplace concepts. Results Participants described trade-offs concerning target users and host, component scope and customization, quality labels, verification, geographic scope, pricing, interoperability, platform launch, risks and market niche. The resulting concepts emphasized a regional startup ecosystem, a research-oriented hybrid marketplace or a global marketplace with stricter entry requirements. Discussion The concepts combine the tensions in different ways and highlight competing priorities in governance, openness, assurance, scalability and early platform growth. Conclusion Stakeholders identified recurring design choices for a DTx software-component marketplace. The concepts provide hypotheses for prototyping and evaluation; the study did not test technical feasibility, market demand, regulatory acceptability or effects on development cost or time.

15
Aggregating data to accelerate personalized therapy in heart failure (ADAPT-HF)

Roeder, C.; Goerg, C.; Talebi, A.; Stevens, L. M.; Scholtens, D. M.; Rasmussen-Torvik, L. P.; Alagna, L. M.; Shah, S. J.; Hall, J. L.; Das, A. K.; Jhund, P. S.; Kao, D. P.

2026-07-16 health informatics 10.64898/2026.07.13.26357501 medRxiv
Top 0.9%
0.3%
Show abstract

Background: Increased public access to data from disparate sources provides opportunities to study and validate predictive and subphenotype models in heterogeneous disease conditions using aggregated individual patient data. Robust, explicit, and transparent harmonization of data elements is critical to ensure interpretability, reproducibility, and generalizability of secondary and retrospective analyses. Methods & Results: We designed and implemented ADAPT (Aggregating Data to Accelerate Personalized Therapy), a scalable framework using multiple software packages (R, SQL, BigQuery) that enables rapid, explicit harmonization of structured data elements from randomized trials and observational studies using a standard spreadsheet interface. User-specified criteria are applied to primary study data to produce harmonized longitudinal datasets comprised of demographics, medical history, quantitative observations, repeated measures, and clinical outcomes. We demonstrate this functionality using 26 clinical studies found in the National Heart, Lung, and Blood Institute BioLINCC resource. We illustrate the scalability of ADAPT to the order of billions of datapoints using administrative clinical data in a cloud-computing platform. We also present examples of collaborators using ADAPT for independent harmonization tasks for secondary analyses and democratization of publicly available data. Conclusion: ADAPT is a disease-agnostic, extensible, and scalable platform to support robust, transparent harmonization of structured research data using interfaces accessible to a variety of researchers regardless of programming ability. It extends FAIR principles beyond research data to also represent harmonization analyses by improving Findability of harmonization decisions, Accessibility of methods to other stakeholders, Interoperability with independent analyses and datasets, and Reusability through efficient implementation in a variety of analysis environments.

16
In Silico Trial Simulation with Artificial Intelligence-Generated Synthetic Control Cohorts Reproduces Results of a Randomized Controlled Trial in Acute Myeloid Leukemia

Kumar Reddy, K.; Hahn, W.; Winter, S.; Roellig, C.; Mueller-Tidow, C.; Serve, H.; Baldus, C. D.; Fransecky, L.; Schliemann, C.; Burchert, A.; Schaefer-Eckart, K.; Kaufmann, M.; Schetelig, J.; Bornhaeuser, M.; Middeke, J. M.; Eckardt, J.-N.

2026-07-16 health informatics 10.64898/2026.07.15.26358123 medRxiv
Top 0.9%
0.3%
Show abstract

Rising costs, slow accrual and molecular substratification of cancers necessitate novel clinical trial designs. We demonstrate that artificial intelligence-generated synthetic patients can replace real controls to reproduce results of the SORAML trial. Using external multimodal data from 1,377 acute myeloid leukemia (AML) patients from previous trials and a real-world registry, we fine-tuned a tabular foundation model to generate synthetic patients, reproducing clinical and genetic features and outcome associations. Synthetic patients were then matched to the original SORAML intervention group using Cox risk scores, replacing the original control and reproducing the original trial result with near-identical median event-free survival (EFS) and treatment effect (original hazard ratio [HR] 0.64, 95%-confidence interval [CI] 0.47-0.87, p=0.004; with synthetic control HR 0.66, 95%-CI 0.48-0.90, p=0.009). Our findings demonstrate that AI-generated synthetic patients can serve as statistically rigorous controls supporting novel trial designs.

17
Statistical Analysis of Pre-War Primary Healthcare Costs in Ukraine: Variations by Location and Ownership and Implications for Financing Reform

Mohamed, A. T.; Kasekamp, K.; Demeshko, O.; Habicht, T.; Murphy, A.; Sadique, Z.

2026-07-16 health policy 10.64898/2026.07.15.26358144 medRxiv
Top 0.9%
0.3%
Show abstract

Strong primary healthcare (PHC) is associated with lower costs and better population health outcomes when supported by appropriate financing. Costing analysis enables evidenced-based decisions for estimating budgets for PHC and defining provider payments. In 2021, a project supported by the World Health Organization was launched in Ukraine to collect cost data from 100 PHC providers. The objective was to assess costs for delivering services within the state-funded benefits package, with the aim of informing tariff-setting, and assessing budget need. This study used statistical analysis on the collected cost data. We applied multivariable linear regression (MLR) to assess variation in cost-per-person across locality (rural vs. urban) and ownership type (public vs. private) of the providers, after adjusting for confounders. The mean (standard deviation) cost-per-person across the sample providers was 45.46 (18.46) USD. MLR analysis showed that rural providers had a higher cost-per-person of 6.70 USD (95% CI: 1.54, 11.85) compared to urban providers, after adjusting for confounding (p=0.011). We also found strong evidence that private providers had a lower cost-per-person of 36.15 USD (95% CI: -41.82,-30.48) compared to public providers, after adjusting for confounding (p<0.001). Although our findings do not capture the impact of the Russian hostile invasion of Ukraine, they still provide valuable insights for policy discussions within Ukraine and for other nations examining PHC financing reforms. Our findings align with international evidence suggesting that rural providers incur higher costs, supporting the need to adjust capitation payments for providers in these areas. Ownership type also affects costs, potentially reflecting differences in quality standards between public and private providers. These differences allow private providers to opportunistically reduce costs by limiting staff numbers and optimizing facility size to maximize profits. To ensure equitable access to high-quality PHC, uniform service delivery standards should be applied to all PHC providers, regardless of ownership type.

18
The Variance-Stabilizing Transformation for the Poisson Rate Ratio: Closed-Form Confidence Intervals

Ng, S.-P.

2026-07-18 epidemiology 10.64898/2026.07.16.26358255 medRxiv
Top 1.0%
0.3%
Show abstract

The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.

19
Developing and Prospectively Validating a Reproducible Graph Representation Specification for Clinical Guideline Algorithms: The Measurement Foundation of the Clinical Guideline Complexity Index

Milani, R. V.; Bober, R. M.

2026-07-20 health informatics 10.64898/2026.07.17.26358358 medRxiv
Top 1%
0.2%
Show abstract

Background. Translating a clinical guideline decision algorithm into a computational graph requires judgment, and unconstrained coding yields divergent graphs; any complexity measure computed from such a graph inherits that variation, so its reproducibility must be demonstrated rather than assumed. Objective. To develop, and prospectively test, an empirical method for making graph extraction reproducible, using the Clinical Guideline Complexity Index (CGCI) and four guideline algorithms as a case study. Methods. We built a Graph Representation Specification (an ontology, a motif catalogue, disambiguation conventions, decomposition rules, a deterministic validator, and a scoring engine) and refined it by error-driven grammar induction: measure inter-coder disagreement, localize its dominant class, induce a single grammar rule, and prospectively test whether that rule improves agreement in the anticipated class. Reproducibility was quantified with a pre-specified, topology-based endpoint (Decision Topology Agreement) rather than edge agreement, which is oversensitive to representational choices that do not affect the score. Two trained coders independently coded the diabetes, dyslipidemia, heart-failure, and hypertension algorithms. Results. A rule induced from the diabetes comorbidity panel (assessment topology) generated a pre-specified prediction that heart-failure figures, sharing the same motif, would converge; on a fresh, independently coded pair they did, with an absolute CGCI difference of approximately one. Decision topology reproduced closely (decision-order agreement at or near 1.00 for three of four guidelines), while breadth counting was rule-sensitive: an explicit modifier-counting rule reduced the largest disagreement from 27 to 4 tokens. Residual disagreement was bounded and localizable to specific, nameable representational choices. Conclusions. Graph-extraction reproducibility can be systematically improved through iterative grammar refinement, and a prospectively derived rule can be confirmed to improve agreement. These results establish the measurement foundation (reliability, not construct validity) for a companion study interpreting CGCI as cognitive load, and the method may apply wherever graphs are extracted from structured source artifacts.

20
CAUSAL-RSV: Causal Analysis of RSV Vaccine Effects in Infants Using Real-World Data

Regan, A. K.; Coates, M. M.; Sullivan, S. G.; Munoz, F. M.; Rowe, S. L.; Avila, C.; Arah, O. A.

2026-07-18 infectious diseases 10.64898/2026.07.16.26356876 medRxiv
Top 1%
0.2%
Show abstract

Respiratory syncytial virus (RSV) contributes to substantial morbidity and mortality in young children each year. In 2023, two new prevention products were licensed and recommended in the United States (US), including a prefusion F protein subunit vaccine (RSVpreF) administered during pregnancy and a long-acting monoclonal antibody (mAb) administered in infants. Although post-licensure real-world studies support the effectiveness of RSVpreF vaccine during pregnancy, existing studies have been conducted in settings where only RSVpreF vaccine is available. The real-world effectiveness of RSVpreF vaccine in settings where both RSVpreF vaccine and mAbs are available is not yet well understood. The goal of this study is to estimate the real-world effectiveness of the RSVpreF vaccine against severe infant RSV by applying causal mediation analysis with receipt of mAbs as a mediating variable. Using a national cohort of mother-infant dyads with the Optum Labs Data Warehouse (OLDW), we will model vaccine and mAb effects in a longitudinal cohort spanning the 2023-24, 2024-25, and 2025-26 RSV seasons. Results will be used to better understand the total effect of RSVpreF vaccination when it is used as one component within a hybrid infant RSV prevention program.